AITopics | xiaohua zhai

Collaborating Authors

xiaohua zhai

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

If you are looking for an answer to the question What is Artificial Intelligence? and you only have a minute, then here's the definition the Association for the Advancement of Artificial Intelligence offers on its home page: "the scientific understanding of the mechanisms underlying thought and intelligent behavior and their embodiment in machines."

However, if you are fortunate enough to have more than a minute, then please get ready to embark upon an exciting journey exploring AI (but beware, it could last a lifetime) …

Assaying Generalization

Neural Information Processing SystemsFeb-19-2026, 00:49:48 GMT

Since proxy invariance indiff approaches data.

artificial intelligence, inneurip, machine learning, (18 more...)

Neural Information Processing Systems

Country:

Europe > Germany > Baden-Württemberg > Tübingen Region > Tübingen (0.04)
Europe > France (0.04)
Europe > Denmark (0.04)
Asia > Middle East > Jordan (0.04)

Technology: Information Technology > Artificial Intelligence > Machine Learning (1.00)

Add feedback

LocCa: Visual Pretraining with Location-aware Captioners

Neural Information Processing SystemsFeb-18-2026, 06:40:27 GMT

Specifically, LocCa employs two tasks, bounding box prediction and location-dependent captioning, conditioned on the image pixel input.

large language model, machine learning, natural language, (22 more...)

Neural Information Processing Systems

Country:

Europe > Switzerland > Zürich > Zürich (0.04)
Europe > Belgium > Flanders > Flemish Brabant > Leuven (0.04)

Genre:

Research Report > Experimental Study (0.93)
Research Report > New Finding (0.67)

Technology:

Information Technology > Sensing and Signal Processing > Image Processing (1.00)
Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
(2 more...)

Add feedback

RevisitingNeuralScalingLaws inLanguageandVision

Neural Information Processing SystemsFeb-10-2026, 15:48:02 GMT

The remarkable progress in deep learning in recent years is largely driven by improvements in scale, where bigger models are trained on larger datasets for longerschedules.

artificial intelligence, deep learning, machine learning, (17 more...)

Neural Information Processing Systems

Country:

Europe > Switzerland > Zürich > Zürich (0.14)
Oceania > Australia > Victoria > Melbourne (0.04)

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.67)

Add feedback

RevisitingNeuralScalingLaws inLanguageandVision

Neural Information Processing SystemsFeb-10-2026, 15:47:58 GMT

The remarkable progress in deep learning in recent years is largely driven by improvements in scale, where bigger models are trained on larger datasets for longerschedules.

artificial intelligence, deep learning, machine learning, (18 more...)

Neural Information Processing Systems

Country:

Europe > Switzerland > Zürich > Zürich (0.14)
Oceania > Australia > Victoria > Melbourne (0.04)

Technology: Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning (0.67)

Add feedback

TowardsOpen-VocabularySemanticSegmentation WithoutSemanticLabels

Neural Information Processing SystemsFeb-8-2026, 02:25:15 GMT

Recently, several studies [11, 12, 7, 8] have pioneered open-vocabulary semantic segmentation without densely-annotated semantic labels.

machine learning, natural language, segmentation, (17 more...)

Neural Information Processing Systems

Country: Asia > Middle East > Israel > Tel Aviv District > Tel Aviv (0.04)

Genre: Research Report (0.46)

Technology:

Information Technology > Artificial Intelligence > Natural Language (0.94)
Information Technology > Artificial Intelligence > Vision (0.69)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks (0.46)

Add feedback

LocCa: Visual Pretraining with Location-aware Captioners Bo Wan 1,3 Michael Tschannen 1 Y ongqin Xian

Neural Information Processing SystemsOct-10-2025, 17:36:55 GMT

Specifically, LocCa employs two tasks, bounding box prediction and location-dependent captioning, conditioned on the image pixel input.

dataset, locca, resolution, (17 more...)

Neural Information Processing Systems

Country:

Europe > Switzerland > Zürich > Zürich (0.04)
Europe > Belgium > Flanders > Flemish Brabant > Leuven (0.04)

Genre:

Research Report > Experimental Study (0.93)
Research Report > New Finding (0.67)

Technology:

Information Technology > Sensing and Signal Processing > Image Processing (1.00)
Information Technology > Artificial Intelligence > Vision (1.00)
Information Technology > Artificial Intelligence > Representation & Reasoning (1.00)
(2 more...)

Add feedback

A Recipe for Improving Remote Sensing VLM Zero Shot Generalization

Barzilai, Aviad, Gigi, Yotam, Helmy, Amr, Silverman, Vered, Refael, Yehonathan, Jaber, Bolous, Shekel, Tomer, Leifman, George, Beryozkin, Genady

arXiv.org Artificial IntelligenceMar-17-2025

Foundation models have had a significant impact across various AI applications, enabling use cases that were previously impossible. Contrastive Visual Language Models (VLMs), in particular, have outperformed other techniques in many tasks. However, their prevalence in remote sensing (RS) is still limited, due to the scarcity of diverse remote-sensing visual-language datasets. In this work we introduce two novel image-caption datasets for training of remote sensing foundation models. The first dataset pairs aerial and satellite imagery with captions generated by Gemini using landmarks extracted from Google Maps. The second dataset utilizes public web images and their corresponding alt-text, filtered for the remote sensing domain, resulting in a diverse dataset with greater breadth in image styles and subject matter. These datasets are used to pre-train the MaMMUT~\citep{kuo2023mammutsimplearchitecturejoint} VLM architecture, resulting in state-of-the-art generalization performance in zero-shot cross-modal retrieval on well-known public benchmarks. Finally, we present our ongoing research to distill image-level knowledge gained in the VLM contrastive training procedure to enhance the model's localization ability. Specifically, we iteratively generate pseudo-labels for image regions based on the model's attention maps and use these labels for further training. To mitigate noisy attention maps and create robust segmentation masks, we introduce a novel attention-pooling mechanism called the Smooth-Attention-Operation.

artificial intelligence, large language model, natural language, (16 more...)

arXiv.org Artificial Intelligence

2503.08722

Genre: Research Report (0.66)

Industry: Energy > Renewable > Geothermal > Geothermal Energy Exploration and Development > Geophysical Analysis & Survey (1.00)

Technology: Information Technology > Artificial Intelligence > Natural Language > Large Language Model (1.00)

Add feedback

Nomic Embed Vision: Expanding the Latent Space

Nussbaum, Zach, Duderstadt, Brandon, Mulyar, Andriy

arXiv.org Artificial IntelligenceJun-6-2024

This technical report describes the training of nomic-embed-vision, a highly performant, open-code, open-weights image embedding model that shares the same latent space as nomic-embed-text. Together, nomic-embed-vision and nomic-embed-text form the first unified latent space to achieve high performance across vision, language, and multimodal tasks.

dataset, encoder, text encoder, (14 more...)

arXiv.org Artificial Intelligence

2406.18587

Genre: Research Report (0.40)

Technology:

Information Technology > Artificial Intelligence > Machine Learning (1.00)
Information Technology > Artificial Intelligence > Natural Language (0.99)

Add feedback

Trends in AI -- June 2022

#artificialintelligenceJun-9-2022, 18:02:08 GMT

Originally published on Towards AI the World's Leading AI and Technology News and Media Company. If you are building an AI-related product or service, we invite you to consider becoming an AI sponsor. At Towards AI, we help scale AI and technology startups. Let us help you unleash your technology to the masses. As we go into June, the AI world doesn't stop and once again the pace of new stories and research was high. The ACL conference was held in the past month in Dublin, being one of the first major conferences to go back in person, which certainly feels like another step forward into normalcy.

image generation, representation, transformer, (16 more...)

#artificialintelligence

Technology:

Information Technology > Artificial Intelligence > Vision (0.96)
Information Technology > Artificial Intelligence > Natural Language > Large Language Model (0.71)
Information Technology > Artificial Intelligence > Machine Learning > Neural Networks > Deep Learning > Generative AI (0.31)

Add feedback

Filters

Collaborating Authors

xiaohua zhai

Information about AI from the News, Publications, and Conferences

Automatic Classification – Tagging and Summarization – Customizable Filtering and Analysis

Assaying Generalization

LocCa: Visual Pretraining with Location-aware Captioners

RevisitingNeuralScalingLaws inLanguageandVision

RevisitingNeuralScalingLaws inLanguageandVision

Near_OOD_with_pre_training (1).pdf

TowardsOpen-VocabularySemanticSegmentation WithoutSemanticLabels

LocCa: Visual Pretraining with Location-aware Captioners Bo Wan 1,3 Michael Tschannen 1 Y ongqin Xian

A Recipe for Improving Remote Sensing VLM Zero Shot Generalization

Nomic Embed Vision: Expanding the Latent Space

Trends in AI -- June 2022